Didn't just buy filter-passers in score order.
- Clustered by ESM-2 embedding (PD-L1/Campaign/PDL1-R1/Run/ESM-001, PD-L1/Campaign/PDL1-R1/Run/CLUST-001) to group overlapping binding modes rather than let one lucky region of sequence space fill the whole shortlist.
- Picked a representative per cluster.
- Dropped anything with severe developability warnings (free cysteines, hydrophobic runs, furin sites) regardless of how it scored computationally.
- Made sure positive and negative controls were included in the physical order, not just the 24 candidates.
The top i_pTM design was not the strongest binder measured at stage K, and that's not a fluke — it's the expected failure mode of this kind of metric. i_pTM and pLDDT are trained to answer "how confident is AlphaFold in this fold," not "how tight is this interaction." A design can earn a very high i_pTM by forming a small, very well-packed, high-confidence contact patch — genuinely real, just not large — while a design ranked several points lower buries more interface area, picks up more interface H-bonds, and ends up with the better KD. Shape complementarity, buried SASA and ΔΔG track affinity somewhat better than i_pTM alone, but none of them replace the measurement. Practical consequence: rank candidates for the physical order by a combination of computed metrics and diversity, not by i_pTM sorted descending, and expect the BLI ranking to reshuffle the computed ranking at least a little every round.